fix(srt): align B200 and MI355X recipe images with master configs / fix(srt):对齐 B200 与 MI355X 配方镜像与主配置 - #3567
chunfangamd wants to merge 9 commits into
Conversation
#3428 ported these AgentX recipes from the legacy scripts as they were before the 09-25 image bumps (#3361, #3362, #3334, #3420), which changed only the master images. Every point of these keys now fails before submission with 'Single-node SRT image: recipe/matrix'. The SGLang v0.5.20 recipes also take the --cuda-graph-max-bs-decode rename that #3362 and #3334 applied to the legacy scripts; v0.5.20 no longer accepts the deprecated --cuda-graph-max-bs alias (sgl-project/sglang#38375). Co-authored-by: Cursor <[email protected]>
single_node.py rejects a point whose recipe container differs from the matrix image only once a GPU job starts, and the multi-node path has no such check: srtctl pulls a literal container missing from the alias map. Check both statically, including identity.container.image. Co-authored-by: Cursor <[email protected]>
Co-authored-by: Cursor <[email protected]>
|
Thanks for the contribution!
中文感谢你的贡献!
|
Co-authored-by: Cursor <[email protected]>
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36521758562 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36521758562 |
There was a problem hiding this comment.
Looks good, straightforward config alignment change.
What was reviewed: verified each of the four bumped model.container values against the corresponding (unchanged) master-config image in amd-master.yaml/nvidia-master.yaml — all match exactly; checked that every cuda-graph-max-bs occurrence in the two touched v0.5.20 sglang recipes was renamed to cuda-graph-max-bs-decode, and confirmed no other recipe in the repo still pairs the v0.5.20 sglang image with the old flag name; confirmed the new test drives real recipe expansion (generate_override_configs via selected_recipes) rather than pinning static config, and that the perf-changelog diff only appends a new entry at the file's tail.
Extended reasoning...
The diff is confined to four recipe YAML image/flag edits, a new pure-Python static test, and an append-only changelog entry — no auth, crypto, or data-exposure surface. I cross-checked all four container values against the master configs directly and confirmed they match, verified the flag rename is complete and scoped correctly, and confirmed the new test exercises real config-selection logic rather than pinning literals; the perf-changelog append preserves history. None of the touched paths fall under a CODEOWNERS-restricted pattern in .github/CODEOWNERS. The pull/XXX placeholder in the changelog PR-link is a known pre-merge fill-in matching existing repo convention, not a functional defect.
This review covers commit 75c7220, which is no longer the latest commit on this pull request; later commits are not covered by it.
Restore the MI325X GLM-5.2 and MI300X MiniMax-M3 recipes and drop the static image test from this PR; they will follow separately. The changelog entry now selects only the B200 and MI355X keys. Co-authored-by: Cursor <[email protected]>
…sistency Co-authored-by: Cursor <[email protected]> # Conflicts: # inferencex-e2e/perf-changelog.yaml
Restore the static check with the MI325X GLM-5.2 and MI300X MiniMax-M3 keys exempt until their recipes are aligned, and run CI Tests when master configs or srt-slurm recipes change so YAML-only PRs are checked before any GPU job. Co-authored-by: Cursor <[email protected]>
|
Compared with #3334's passing run (same
|
|
/stage-results 36521758562 |
|
@chunfangamd staged run 36521758562: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-09-29~r36521758562 This run remains available across future |
|
/reuse-sweep-run 36521758562 |
Review showed the pytest scan cannot gate run-sweep, compares image sets per recipe instead of each matrix point's selected variant, misses EVAL_CONFIG_FILE, and pins checked-in config against the AGENTS.md test rules. A separate PR replaces it with a validator that the sweep calls before dispatch. Co-authored-by: Cursor <[email protected]>
Description
Two single-node AgentX srt-slurm recipes name a container that differs from their master-config
image.infx/srt_slurm/single_node.pytherefore rejects every point of these keys before Slurm submission (Single-node SRT image: recipe/matrix ...), the same failure #3446 hit in its first sweep.Cause. #3428 ported these recipes from the legacy scripts as they were before the image bumps in #3420 and #3334, which merged about ten hours earlier. Those bumps changed only the master images, which was correct while the keys still ran the legacy scripts. The PRs touched different files, so git saw no conflict, and no check compares the two copies before a GPU job starts.
image(unchanged)model.containeronmaindsv41flash-fp4-mi355x-vllm-agentic-dsparkvllm/vllm-openai-rocm:nightly-rocm100-29468dde…vllm/vllm-openai-rocm:nightly-rocm100-7f1a5398…dsv4-fp4-b200-sglang-agentic-hicache-mtplmsysorg/sglang:v0.5.20-cu130lmsysorg/sglang:v0.5.19-cu130Changes
model.containerto the master image. The B200 SGLang v0.5.20 recipe also renamescuda-graph-max-bstocuda-graph-max-bs-decodein all 12 variants, as [Klaud Cold] Update dsv4-fp4-b200-sglang-agentic-hicache-mtp SGLang image to v0.5.20-cu130 / 将 dsv4-fp4-b200-sglang-agentic-hicache-mtp 的 SGLang 镜像更新至 v0.5.20-cu130 #3334 did in the legacy script. SGLang v0.5.20 no longer accepts the deprecated alias ([Config] Retire get_global_server_args, and clear the deprecated flags that have a replacement sgl-project/sglang#38375), so an image-only fix would fail at server startup. No other serving flag or sweep point changes.perf-changelog.yaml: one entry for the two keys so the sweep re-validates them on the native srt-slurm path.Earlier commits also added a static recipe/master image test. Review showed that a pytest scan cannot gate
run-sweep, compares image sets per recipe instead of each matrix point's selected variant, missesEVAL_CONFIG_FILE, and pins checked-in config against theAGENTS.mdtest rules. It is removed here and replaced by #3624, a validator thatrun-sweepcalls before dispatch.Validation
mi355x-amdsandb200-nscale. Earlier attempts failed on infrastructure, not on the recipes:g15,g18andg37kept GPU memory with no owning process after other jobs crashed or were cancelled./reuse-sweep-run 36521758562was accepted for this PR.Notes for reviewers
main, a sweep that selects either key fails its preflight until these recipes are onmaintoo.dsv41flashcontainer line, so whichever merges second resolves a one-line conflict.glm5.2-fp8-mi325x-sglang-agentic-mtpstill has the same drift onmainand is left for a separate change.minimaxm3-fp8-mi300x-vllm-agentic-mtpwas aligned onmainby feat(agentx): run the MiniMax-M3 MI300X LMCache point on the lmcache-server service #3545.@SemiAnalysisAI/core.AI model disclosure
Related Issue
No issue. Follow-up: #3624. Related: #3428, #3446, #3555.
Type of Change
Checklist
inferencex-e2e/perf-changelog.yamland have not edited historical entriesOWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.中文
改动说明
两个单节点 AgentX srt-slurm 配方的 container 与主配置
image不一致,导致infx/srt_slurm/single_node.py在提交 Slurm 之前拒绝这些 key 的每一个点(Single-node SRT image: recipe/matrix ...),与 #3446 第一次 sweep 遇到的失败相同。原因: #3428 移植这些配方时,依据的是 #3420 和 #3334 升级镜像之前的旧脚本,而这两个升级约在 #3428 合入前十小时已经合入。升级 PR 只改了主配置镜像,这在这些 key 仍运行旧脚本时是正确的。两边改的是不同文件,git 没有冲突,而在 GPU 任务开始之前也没有任何检查比较两份拷贝。上表列出了两个 key 的主配置镜像(未改动)和
main上配方的旧镜像。改动:
model.container改为主配置镜像。B200 的 SGLang v0.5.20 配方同时在全部 12 个 variant 中将cuda-graph-max-bs改为cuda-graph-max-bs-decode,与 [Klaud Cold] Update dsv4-fp4-b200-sglang-agentic-hicache-mtp SGLang image to v0.5.20-cu130 / 将 dsv4-fp4-b200-sglang-agentic-hicache-mtp 的 SGLang 镜像更新至 v0.5.20-cu130 #3334 对旧脚本的修改一致。SGLang v0.5.20 已移除该弃用别名([Config] Retire get_global_server_args, and clear the deprecated flags that have a replacement sgl-project/sglang#38375),只改镜像会导致服务启动失败。其余服务参数和 sweep 点均不变。perf-changelog.yaml: 为这两个 key 追加一条记录,让 sweep 在新的 srt-slurm 路径上重新验证。之前的 commit 还加了一个配方与主配置镜像的静态测试。Review 指出:pytest 扫描无法拦住
run-sweep;它按配方比较镜像集合,而不是每个 matrix 点实际选中的 variant;没有检查EVAL_CONFIG_FILE;并且在测试里固定了 checked-in 配置,违反AGENTS.md的测试规则。因此本 PR 移除该测试,改由 #3624 提供一个在run-sweep派发前调用的 validator。验证:
mi355x-amds和b200-nscale上 28 个吞吐点和 2 个 AgentX eval 点。之前几次失败都是基础设施问题,与配方无关:g15、g18、g37在其他任务崩溃或被取消后,残留了没有归属进程的显存。/reuse-sweep-run 36521758562已被接受。审阅注意事项:
main之后,选中这两个 key 的 sweep 会在预检时失败,直到本 PR 的配方也进入main。dsv41flashcontainer,后合入的一方需要解决一行冲突。glm5.2-fp8-mi325x-sglang-agentic-mtp在main上仍有同样的不一致,留待单独修复;minimaxm3-fp8-mi300x-vllm-agentic-mtp已由 feat(agentx): run the MiniMax-M3 MI300X LMCache point on the lmcache-server service #3545 在main上对齐。@SemiAnalysisAI/core。AI 模型使用说明
关联 issue
无。后续 PR:#3624。相关 PR:#3428、#3446、#3555。
改动类型
Bug 修复、配置修改。